Skip to content

docs(openspec): chain-restructuring proposal (reject complex one-liners with guidance) - #17

Open
preved911 wants to merge 9 commits into
mainfrom
openspec/chain-restructuring
Open

preved911 wants to merge 9 commits into
mainfrom
openspec/chain-restructuring

Conversation

@preved911

@preved911 preved911 commented Sep 3, 2026 •

Copy link
Copy Markdown
Owner

Summary

OpenSpec change proposal for chain-restructuring — deterministic steering: the plugin rejects not-allowed complex one-liners with an actionable error, so the agent re-issues them as separate commands or a multi-line script (one command per line), each checked individually. Allowed chains stay allowed — restructuring targets exactly the commands a human must review. Proposal only; implementation follows after review.

Why

Not-allowed multi-step commands surface to the human as an unreadable one-liner in the permission dialog, and the agent has no incentive to write readable commands. AGENTS.md instructions are soft (probabilistic, no verification loop).

Verified mechanism (from opencode API research)

  • permission.ask output carries only { status } — no reason field; per issue [FEATURE]: Wire the permission.ask plugin hook anomalyco/opencode#19469 the hook isn't even triggered by the engine in current source. Unusable for steering.
  • tool.execute.before + throw works: thrown errors become tool results with resultType: "error" and the error text reaches the model (official docs pattern). The model retries in compliant form — a closed deterministic loop.

Proposed config — separate plugin file

opencode-bash-guard.jsonc in the opencode config dirs (global ~/.config/opencode/, project .opencode/), JSONC with comments, deep-merged project over global. Permission actions stay in opencode.json; plugin behavior tuning lives here.

{
  // Reject complex one-liners and ask the agent to restructure them
  "restructure": {
    "enabled": false,     // default false — zero behavior change
    "max_segments": 3,    // max commands in a single-line chain
    "max_depth": 2        // max $()/backtick/meta-command nesting
  }
}

Key design decisions

  1. Scope: ask-resolving chains only — allowed = allowed (pass untouched); deny = forbidden regardless of format (unchanged); parse errors = fail-closed deny (unchanged). Under the "*": "ask" prerequisite, "not allowed" ⇔ ask in practice
  2. One-liner targeting: segment limit applies to single-line commands only; multi-line scripts (the compliant form) are exempt from the segment limit but still depth-checked — a compliant re-issue always exists, so the retry loop cannot deadlock
  3. Metrics: segment count (existing parseChain) + max substitution nesting depth (new — catches echo $(echo $(...)) obfuscation)
  4. Reject with guidance: error message includes actual counts + both compliant forms; nothing executes, no dialog
  5. No bypass: multi-line re-issues parse per line (newlines are already segment separators), separate calls are single segments — everything still permission-checked
  6. README correction: removes the stale "multi-segment → ask (defense-in-depth)" claim that contradicts implementation/spec
  7. AGENTS.md snippet documented as complementary soft layer (reduces rejection frequency; plugin guarantees enforcement)

Artifacts

  • proposal.md — what & why
  • design.md — 8 decisions + risks (review decision recorded: restructure applies only to not-allowed commands)
  • specs/chain-restructuring/spec.md — 6 requirements, 26 scenarios
  • tasks.md — 6 sections, 29 tasks (incl. new config-loader module + jsonc-parser dep)

Out of scope (flagged)

permission.ask may never fire in current opencode (anomalyco/opencode#19469) — the deny path of this plugin may not hard-block. Needs separate verification + possible fix change.

Next steps

  • Review proposal
  • Run /opsx-apply to implement

Ultraworked with Sisyphus

preved911 and others added 3 commits September 3, 2026 14:14
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

👀 AI Code Review

Something went wrong: <urlopen error [Errno -2] Name or service not known>


Powered by GPT-4o via GitHub Models

preved911 and others added 2 commits September 4, 2026 14:30
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

👀 AI Code Review

Something went wrong: <urlopen error [Errno -2] Name or service not known>


Powered by GPT-4o via GitHub Models

@preved911

Copy link
Copy Markdown
Owner Author

There is nothing about long python/perl/nodejs scripts execution and parsing. This one should be processed too. Inline scripts should be pretty formatted

@preved911

Copy link
Copy Markdown
Owner Author

Think about jev (https://typesafe.ai/blog/introducing-system-one-models-and-jev) usage for complicated commands detection

@preved911

Copy link
Copy Markdown
Owner Author

Think about AGENTS.md like instructions extension by plugin, without additional manual user steps

- `allow` → untouched, always.
- `deny` → existing deny flow (wrap + stored deny). Restructuring a forbidden action is meaningless — the format is not the problem.
- Parse errors → existing fail-closed deny, unaffected.
- `null` (no plugin opinion) → untouched. Under the documented prerequisite (`"*": "ask"`), every uncovered segment matches the catch-all, so any chain that would reach a human resolves to `ask` — "not allowed" ⇔ `ask` in practice. Without the catch-all the plugin has no opinion and does not intervene.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Default permissions settings must be taken into account

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 01a5386 — decision 3 now records how default permission settings are accounted for.

Concretely (code-backed, parseConfig in src/config.ts): the entire plugin disables itself when permission.bash is absent (opencode's default state — no user permission config) or resolves to allow (flat "bash": "allow" or object "*": "allow"). So under default permissions neither the chain guard nor restructure runs — no surprising prompts for users who never opted in. restructure can only fire when the user actually configured bash permissions, and under the documented "*": "ask" prerequisite "not allowed ⇔ ask" holds exactly; a config without the catch-all leaves uncovered segments at null (no opinion → untouched).

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the entire plugin disables itself when permission.bash is absent (opencode's default state — no user permission config) or resolves to allow (flat "bash": "allow" or object "*": "allow")

Only one of these may happen by default. We need to know what is opencode's default behaviour, when permssions settings not exist

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified — you're right that only one default exists, and it's now recorded precisely (design.md decision 3, 8ee8e5f).

OpenCode's documented defaults are permissive: with no permission settings, bash defaults to "allow" — commands run without any prompt. Only external_directory and doom_loop default to "ask" (Permissions → Defaults).

Consequence for the plugin, now stated exactly in the design: in the default state there is no effective bash rule to resolve against, and parseConfig (src/config.ts) disables the entire plugin (enabled = false when permission.bash is absent, a flat "allow", or "*": "allow" — all resolve to the same inert state). So out of the box: native behavior = allow-all, plugin inert, no prompts introduced. restructure can only ever fire in a user-configured permission setup, and with the documented "*": "ask" prerequisite "not allowed ⇔ ask" holds exactly.


Flow in `beforeExecute`: parse → resolve chain (existing) → if action is `ask` AND `restructure.enabled` AND limits exceeded → **throw** (replaces wrap+store for that call). Otherwise existing flows verbatim.

4. **One-liner targeting: segment limit applies to single-line commands only**

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure about it. If we have two lines command with a lot of segments, it still should be splitted

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed — replaced the multi-line exemption with a per-line check (01a5386, design decision 4 + spec "Detect command complexity").

The segment limit now applies to every line of every command shape: a single-line command is one line, a multi-line script gets each line checked individually. A 2-line script with many segments per line is rejected, and the message is line-scoped (line 2: 5 chained commands, nesting depth 2) so the model knows exactly which line to fix.

Deadlock-freedom is preserved: the compliant form (one command per line) has every line at 1 segment, so a compliant re-issue always exists.

- Command contains no newline → both `max_segments` and `max_depth` checks apply.
- Command contains a newline (multi-line script) → **exempt from the segment limit**; depth check still applies.

Rationale: the compliant form ("multi-line, one command per line") must never violate the segment limit, otherwise the retry loop can deadlock (a 5-step task re-issued as 5 lines would still exceed `max_segments: 3` and be rejected forever). Multi-line IS the readable form the feature asks for. Deep nesting remains unreadable in any shape, so `max_depth` applies to both forms. Known edge, accepted for v1: a multi-line script with long `&&`-chains inside individual lines passes the segment check (ask still gates uncovered segments; a per-line segment check is possible future refinement).

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Known edge, accepted for v1: a multi-line script with long &&-chains inside individual lines passes the segment check (ask still gates uncovered segments; a per-line segment check is possible future refinement).

If whole command will be constructed on && and || operands, it wouldn't be splitted? I'm afraid that it is frequent case

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right, and it is the frequent case — fixed by the same change as the thread above (01a5386): the "accepted v1 edge" is gone. The segment limit now applies per line to every command shape, so a "multi-line" command built on &&/|| chains is checked line-by-line; any line over max_segments on an ask-resolving chain is rejected with guidance to split it further (design decision 4, risks updated, spec scenario "Multi-line script with long per-line chains — rejected" + task 4.4 added).

The always-compliant form remains one command per line (1 segment per line), so the retry loop cannot deadlock while the &&-wrapped blobs can no longer slip through the multi-line shape.

…faults + review notes

- Decision 4: segment limit applies per line to every command shape (multi-line
  is not an exemption); rejection message names the offending line
- Decision 9 + spec/tasks: interpreter inline scripts (python -c, node -e,
  perl -e, ...) complexity-checked via statement count; pretty formatting
  (one statement per line) as the compliant form
- Decision 3: record how default permission settings are accounted for
  (plugin self-disables when bash permission config is absent or allow-all)
- Future directions: plugin-injected instructions, TypeSafe Jev classifier

Addresses review threads and owner comments on PR #17.

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
@github-actions

Copy link
Copy Markdown

👀 AI Code Review

Something went wrong: <urlopen error [Errno -2] Name or service not known>


Powered by GPT-4o via GitHub Models

@preved911

Copy link
Copy Markdown
Owner Author

Added to scope in 01a5386 (design decision 9, new spec requirement "Complexity-check interpreter inline scripts", tasks 2.5/3.1–3.4/4.4).

What's covered: python/python3 -c, perl -e, node -e/--eval, ruby -e, php -r, and heredoc-scripted interpreters. These carry whole programs inside a single shell segment, so shell-level segment/depth metrics never see a 40-statement node -e blob. When restructure is enabled, the plugin adds an inline-script statement count (split on ; and newlines) checked against the same max_segments threshold.

Pretty formatting as the compliant form: over-threshold inline scripts are rejected with interpreter-specific guidance — re-issue with one statement per line inside the quoted script (python -c, node -e, perl -e all accept multi-line script strings) or move the script to a file.

Deliberately a deterministic heuristic (no per-language parsers in v1); the miscount risk (semicolons inside strings) is documented under Risks — it errs toward rejection of unreadable blobs, and per-language AST parsing is a possible refinement.

@preved911

Copy link
Copy Markdown
Owner Author

Researched the post — Jev is a good conceptual fit for exactly the fuzzy part of this problem.

Where it fits: our detection is AST-threshold based (max_segments, max_depth, statement counts) — deterministic but brittle at the margins. "Is this command too complex/unreadable to review?" is a fuzzy judgment, and Jev's shape (unstructured command in → typed calibrated decision out, ~70–500 ms, cheap, schema outputs that can't hallucinate types) matches a "score this command" detector well.

Why not in v1 (recorded under design.md → Future Directions in 01a5386):

  • This change's whole point is deterministic steering — a closed loop with no external dependency. A guard that phones home adds a trust/privacy surface (every command would leave the machine) and an availability dependency on a security path.
  • Jev is early-access; no stable contract to code against yet.

Proposed follow-up: an optional, off-by-default detector backend ("detector": "ast" | "jev" in opencode-bash-guard.jsonc) where the deterministic AST path remains the source of truth and the fallback when the service is unavailable. Worth a spike once early access is in hand — the calibrated probabilities could also drive smarter UX (e.g., only reject when complexity confidence is high, let borderline ones through to the human dialog).

…nd in #22

- Non-Goals: no learned/external classifier in v1
- Future Directions: Jev bullet marked "decided: NOT in v1" with rationale
  (deterministic-only detection, no privacy surface, no availability coupling)
  and linked to #22 for the follow-up detector backend
@github-actions

Copy link
Copy Markdown

👀 AI Code Review

Something went wrong: <urlopen error [Errno -2] Name or service not known>


Powered by GPT-4o via GitHub Models

@github-actions

Copy link
Copy Markdown

👀 AI Code Review

Something went wrong: <urlopen error [Errno -2] Name or service not known>


Powered by GPT-4o via GitHub Models

@preved911

Copy link
Copy Markdown
Owner Author

Re: plugin-managed AGENTS.md instructions without manual user steps — yes, this is possible today.

OpenCode's plugin API has an experimental hook "experimental.chat.system.transform". Its signature is (input: { sessionID?: string; model: Model }, output: { system: string[] }) => Promise<void> (plugin types). The opencode core calls it on every LLM request after assembling the base system prompt (request.ts), so a plugin can append guidance by mutating output.system. That would let opencode-bash-guard ship its "one command per line" guidance itself, with no user AGENTS.md step.

A few related facts:

  • AGENTS.md itself is auto-loaded by OpenCode V2: root/global files are discovered automatically; nested files are loaded as the agent reads files in those directories (docs). So the snippet in decision 7 already works without extra config for root-level rules.
  • The instructions array in opencode.json exists in the V2 schema but the docs note V2 does not currently resolve it.
  • There's existing precedent: @qforge/opencode-agents-explorer is a plugin that auto-injects folder-level AGENTS.md files via hooks.

Because the hook is experimental, I kept the AGENTS.md snippet as the v1 soft layer and recorded auto-injection as a future direction in design.md. We can add an opt-in plugin flag later if you want it.

…llow"

Permissions docs (Permissions → Defaults): with no permission settings, most
permissions default to "allow"; only external_directory and doom_loop default
to "ask". So the out-of-box state is allow-all bash; parseConfig self-disables
the plugin there (absent bash config, flat "allow", or "*": "allow" all resolve
to the same inert state). restructure can only fire in a user-configured
permission setup.

Clarifies review thread r4130311247 on PR #17.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant